Papers with knowledge distillation approaches
On Knowledge distillation from complex networks for response prediction (N19-1)
Copied to clipboard
| Challenge: | Recent advances in Question Answering have led to the development of very complex models . however, these models are expensive in space and time and require limited resources . |
| Approach: | They propose to use simple models which learn to emulate characteristics of a teacher network . they use a 12GB Tesla K80 GPU to restrict the maximum length of the input document . |
| Outcome: | The proposed model can perform better on a Holl-E dialog dataset. |
BERT-of-Theseus: Compressing BERT by Progressive Module Replacing (2020.emnlp-main)
Copied to clipboard
| Challenge: | a novel approach to compress neural networks by progressive module replacement is proposed . a number of techniques have been proposed to compress pretraining and fine-tuning models . |
| Approach: | They propose a model compression approach that divides BERT into modules and builds their compact substitutes. |
| Outcome: | The proposed approach outperforms existing knowledge distillation approaches on GLUE benchmark . it is based on a model that divides the original BERT into several modules and builds their substitutes . |